NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Weaver: Efficient Coflow Scheduling in Heterogeneous Parallel Network

https://doi.org/10.1109/IPDPS47924.2020.00113

Huang, Xin Sunny; Xia, Yiting; Ng, T. S. (May 2020, 4th IEEE International Parallel & Distributed Processing Symposium (IPDPS 2020), New Orleans, LA)

Full Text Available
Weaver: Efficient Coflow Scheduling in Heterogeneous Parallel Networks

Huang, Xin Sunny; Xia, Yiting; Ng, T. S. (May 2020, 34th IEEE International Parallel & Distributed Processing Symposium (IPDPS 2020), New Orleans, LA, May 2020)

Full Text Available
Green, Yellow, Yield: End-Host Traffic Scheduling for Distributed Deep Learning with TensorLights

Huang, Xin Sunny; Chen, Ang; Ng, T. S. (May 2019, 5th IEEE International Workshop on High-Performance Big Data and Cloud Computing (HPBDC 2019))

Full Text Available
Exploiting Inter-Flow Relationship for Coflow Placement in Datacenters

https://doi.org/10.1145/3106989.3107004

Huang, Xin Sunny; Ng, T. S. (January 2017, APNet'17 Proceedings of the First Asia-Pacific Workshop on Networking)

A crucial challenge for data-parallel clusters is achieving high application-level communication efficiency for structured traffic flows (a.k.a. Coflows) from distributed data processing applications. A range of recent works focus on designing network scheduling algorithms with predetermined Coflow placement, i.e. the endpoints of subflows within a Coflow are preset. However, the underlying Coflow placement problem and its decisive impact on scheduling efficiency have long been overlooked. It is hard to find good placements for Coflows. At the intra-Coflow level, constituent flows are related and therefore their placement decisions are dependent. Thus, strategies extended from flow-by-flow placement is sub-optimal due to negligence of the inter-flow relationship in a Coflow. At the inter-Coflow level, placing a new Coflow may introduce contentions with existing Coflows, which changes communication efficiency. This paper is the first to study the Coflow placement problem with careful considerations of the inter-flow relationship in Coflows. We formulate the Coflow placement problem and propose a Coflow placement algorithm. Under realistic traffic in various settings, our algorithm reduces the average completion time for Coflows by up to 26%.
more » « less
Full Text Available
Stop Rerouting!: Enabling ShareBackup for Failure Recovery in Data Center Networks

https://doi.org/10.1145/3152434.3152452

Xia, Yiting; Huang, Xin Sunny; Ng, T. S. (January 2017, HotNets-XVI Proceedings of the 16th ACM Workshop on Hot Topics in Networks)

This paper introduces sharable backup as a novel solution to failure recovery in data center networks. It allows the entire network to share a small pool of backup devices. This proposal is grounded in three key observations. First, the traditional rerouting-based failure recovery is ineffective, because bandwidth loss from failures degrades application performance drastically. Therefore, failed devices should be replaced to restore bandwidth. Second, failures in data centers are rare but destructive [11], so it is desirable to seek cost-effective backup options. Third, the emergence of configurable data center network architectures promises feasibility of bringing backup devices online dynamically. We design the ShareBackup prototype architecture to realize this idea. Compared to rerouting-based solutions, ShareBackup provides more bandwidth with short path length at low cost.
more » « less
Full Text Available
A Tale of Two Topologies: Exploring Convertible Data Center Network Architectures with Flat-tree

https://doi.org/10.1145/3098822.3098837

Xia, Yiting; Sun, Xiaoye Steven; Dzinamarira, Simbarashe; Wu, Dingming; Huang, Xin Sunny; Ng, T. S. (January 2017, ACM SIGCOMM'17)

This paper promotes convertible data center network architectures, which can dynamically change the network topology to combine the benefits of multiple architectures. We propose the flat-tree prototype architecture as the first step to realize this concept. Flat-tree can be implemented as a Clos network and later be converted to approximate random graphs of different sizes, thus achieving both Clos-like implementation simplicity and random-graph-like transmission performance. We present the detailed design for the network architecture and the control system. Simulations using real data center traffic traces show that flat-tree is able to optimize various workloads with different topology options. We implement an example flat-tree network on a 20-switch 24-server testbed. The traffic reaches the maximal throughput in 2.5s after a topology change, proving the feasibility of converting topology at run time. The network core bandwidth is increased by 27.6% just by converting the topology from Clos to approximaterandom graph. This improvement can be translated into acceleration of applications as we observe reduced communication time in Spark and Hadoop jobs.
more » « less
Full Text Available

Search for: All records